Skip to content

feat: separate physical planning from deployment and share DAG execution - #462

Open
zzylol wants to merge 90 commits into
mainfrom
feat/shared-physical-operators
Open

zzylol wants to merge 90 commits into
mainfrom
feat/shared-physical-operators

Conversation

@zzylol

@zzylol zzylol commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Dataset-bound SDS alignment

Dataset-bound semantic export now carries an explicit logical dataset (namespace, dataset) in semantic fragment version 2. Different datasets cannot share a definition merely because their metric expressions match. The existing unbound version 1 exporter remains structural; Backend uses the bound exporter for persisted outputs. Semantic identity and serialization regression tests pass.

Before this PR

ASAPPlanner selected logical computations, while deployments implemented physical operators and execution separately. There was no shared library to keep Planner semantics, operator behavior and shared-producer execution consistent.

After this PR

Add asap-physical-operators: native relational and summary operators, a physical compiler, and a runtime for one DAG execution.

Logical Post-ASAP candidate + maintenance requirements
    → Planner physical compilation
    → precompute/query Physical DAGs + typed boundaries
    → deployment selection and source/state binding
    → shared executor + RunContext

Planner owns operators, dependencies, sharing and materialization frontiers. Backend owns deployment feasibility, runtime-statistics/resource/ERP costing, storage bindings and scheduling. Build, merge and readout are reusable operators, not fixed deployment phases. See the design.

The library includes:

  • Reader-independent physical compilation and versioned physical-program recovery; binding does not repeat logical lowering.
  • Bounded frontier enumeration and comparable workload-cost selection, including Rate → grouped Sum placement candidates.
  • Shared producers, isolated runs, bounded backpressure, cancellation and tracked memory budgets.
  • Relational/scalar execution, summary build/merge/readout, and automatic bounded temporal KLL pane compilation.
  • Float64 weighted CMS/CountSketch heaps and versioned native summary payloads; sketch algorithms remain in asap_sketchlib.
  • An explicit full-label source-row representation for open PromQL series. Rate/heap execution preserves labels not named in the query. Native current-series snapshots apply replacement, expiry and stale markers before heap updates, so historical gauge values do not accumulate.

Validation

Operator, physical-DAG and Planner-to-execution tests cover schemas, errors, resources, independent runs, shared dependencies and serialization. Latest physical-package run passed 320 tests including its documentation test; two additional identity/resource tests and the final weighted/current-series suites also pass. Strict all-target Clippy passes.

The weighted binding suite passes eight tests, including both grouped Rate/Sum placements and raw samples through Rate and both heap families. Coverage includes resets, zero scores, shifted windows and unreferenced labels. Current-series tests cover decreases, expiry, stale markers, conflicting timestamps and recovered physical programs. These are shared-library tests, not deployment E2E.

Integration limits

Backend #761 installs and costs native physical candidates for SQL, spatial TopK, Rate TopK and grouped Rate/Sum. Certified deployment tests execute query-time and precomputed CMS/CountSketch heaps and grouped Sum through durable SDS, HTTP and restart. Precomputed aggregates require finite complete-input closure; overlapping full windows preserve the query cadence. #728 now retains 10 selected plans and 34 compiled candidate plans, including both Sum placements, for human review. Other PromQL shapes retain documented Backend adapters; joint workload candidate search remains bounded rather than exhaustive.

General scalar-output persistence, distributed execution, sharding and spill remain outside the established deployment path. Ad-hoc SDS discovery is deferred. No new manual deployment verification or human plan approval is claimed. The previously recorded Level 3 performance gate remains failed.

@zzylol
zzylol marked this pull request as ready for review September 24, 2026 13:13
@zzylol
zzylol force-pushed the feat/shared-physical-operators branch from 50a3972 to 9ceeba7 Compare September 24, 2026 14:28
@zzylol
zzylol force-pushed the refactor/membership-subgraph branch from ad95c6d to 12ccae2 Compare September 24, 2026 14:28
@zzylol
zzylol force-pushed the refactor/membership-subgraph branch from 12ccae2 to 0afd38a Compare September 24, 2026 14:49
@zzylol
zzylol force-pushed the feat/shared-physical-operators branch 2 times, most recently from 2b6171a to 9463710 Compare September 24, 2026 15:00
@zzylol
zzylol force-pushed the refactor/membership-subgraph branch 2 times, most recently from 36e7da7 to da6ecbf Compare September 24, 2026 15:08
@zzylol
zzylol force-pushed the feat/shared-physical-operators branch 3 times, most recently from 7c45150 to d0cbd11 Compare September 24, 2026 16:33
@zzylol
zzylol changed the base branch from refactor/membership-subgraph to main September 24, 2026 16:34
@zzylol zzylol assigned GordonYuanyc and unassigned GordonYuanyc Sep 24, 2026
@zzylol
zzylol requested a review from GordonYuanyc September 24, 2026 16:37
@milindsrivastava1997

milindsrivastava1997 commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

High-level questions @zzylol:

  • What is the aim of this PR? Is it to create a separate library of physical operators? Is it to create a runtime for those physical operators? Both?
  • What are the E2E behavior to be tested here? Are there tests for this? Have these been verified manually?
  • What is the before vs after E2E effect of this PR?
  • Given this PR, what is the expected role of asap-fusion and ASAPQuery?

@zzylol

zzylol commented Sep 24, 2026

Copy link
Copy Markdown
Contributor Author

High-level questions @zzylol:

  • What is the aim of this PR? Is it to create a separate library of physical operators? Is it to create a runtime for those physical operators? Both?

Both.

  • What are the E2E behavior to be tested here? Are there tests for this? Have these been verified manually?

There are two levels:

  1. Each physical operator produces correct results and handles its supported types, edge cases, errors and resource constraints.
  2. Given a physical DAG, the operators execute together correctly, including dependencies, shared producers and independent execution runs.

Both have automated coverage: operator tests and DAG integration tests. These were run locally. No separate manual deployment-level verification has been performed.

The physical DAG (https://github.com/ProjectASAP/ASAPPlanner/blob/2216fb9/crates/asap-physical-operators/src/plan/mod.rs) comes from the Planner’s post-ASAP plan, followed by binding to concrete operators:

SQL / PromQL
↓ frontend
QueryExpr
↓ ASAP planning and summary selection
SummaryExpr
↓ compile_executable_dag(...)
ExecutableDag
↓ binding::bind(...) + deployment sources
PhysicalDag
↓ execute(..., RunContext)
Results

There are two distinct DAGs:

  • ExecutableDag describes the selected computation, schemas and shared dependencies. Despite its name, it does not contain running physical operators.
  • PhysicalDag contains concrete implementations from the shared library. The binder constructs it from ExecutableDag; deployments provide sources or stored-state frontiers and select execution roots.

So deployments reuse Planner-generated plans plus the shared binding/execution infrastructure. They do not need to construct physical DAGs manually, although low-level tests can.

This also adds a third testing requirement: verify that Planner-generated DAGs bind correctly, beyond testing individual operators and manually assembled DAGs.

  • What is the before vs after E2E effect of this PR?

Before: There was no shared physical-operator library for deployments to reuse.

After: Deployments can reuse the same physical-operator implementations and physical DAG execution infrastructure. This also provides a common place to implement future physical execution optimizations, such as parallelism and sharding. Those optimizations are not implemented in this PR.

  • Given this PR, what is the expected role of asap-fusion and ASAPQuery?

asap-fusion and asapquery will be the deployment, which refer/import the physical operator lib here, and implement their own runtimes, e.g., the precompute engine with a DAG / subDAG, query engine. Putting the phsycial opereator lib together with ASAPPlanner is as Yancheng said the benefits of

If we consider organizing the physical operator library inside the ASAPPlanner repo, it has two benefits:

when a new IR node is added, we just need one PR for changing both the logical DAG and the physical operator implementation.
It's easier for testing because some APIs are private and it can broaden the testing coverage because logical ASAP IR and physical operators should be consistent.

zzylol and others added 15 commits September 29, 2026 17:49
The design doc called one half of a PhysicalCandidate the "Maintenance
Physical DAG" while the struct calls it `precompute`, and used both words
elsewhere in the same document. The query half had one name throughout.

The crates already keep these apart by layer: asap-aware-mapping owns the
Summary Maintenance Lifecycle, and asap-physical-operators owns the physical
DAGs a lifecycle compiles to. Borrowing the lifecycle's word for a physical
object collapses that distinction, and consumers inherit the fork: the backend
currently carries both vocabularies for the same thing, with a type named
InstalledPostAsapDag whose own validator is called validate_maintenance.

Use `precompute` for the DAG half, matching PhysicalCandidate, and say once
where the halves are introduced that `maintenance` remains the lifecycle's
word. Section 2 and every lifecycle reference are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two doc comments still called one half of a PhysicalCandidate "maintenance"
while the struct field is `precompute`, and one of them sat directly above the
field it described. The identifiers were already right: TemporalPaneMaintenance
does hold lifecycle requirements, and asap-aware-mapping's
SummaryMaintenanceDagExport does export the lifecycle plan, so neither moves.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The trait doc above it now reads "query-time and precompute-time
accumulation"; the method doc two lines down still said maintenance for the
same input.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow #470: the physical compiler consumes the logical Post-ASAP DAG.
The design doc now names the Pre-ASAP and Post-ASAP DAGs as the two
logical layers and records that placement ownership between per-node
phases and frontier enumeration is still open.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the feat/shared-physical-operators branch from 2358b90 to 76fbbf1 Compare September 29, 2026 17:58
zzylol added a commit that referenced this pull request Sep 29, 2026
Logical-planning half of #462; no physical crate.

- Rate heap Top-K and grouped Rate/Sum summary placements (fixed-window
  and query-time Rate candidates, current-series/counter-value inputs)
- Exact-candidate admission: CostModel::summary_support_evidence hook,
  maintained exact values admitted alongside approximate Top-K/Count,
  division accepts exact guarantees
- Candidate-root enumeration (enumerate_candidate_dags[_for_root],
  assemble_selected_query) with budgeted, non-partial inventory
- PromQL series identity: reserved $promql_series_identity column so
  closed schemas can still be recognized as PromQL populations; direct
  rate/increase topk keeps counter-value ranking intent
- Per-series reduction resolves the sample column instead of guessing
- HLL confidence: standalone hll_confidence module dropped; it duplicated
  ClassicHllConfidence already on main in accuracy/estimators/hll.rs (#460)
- Update planner Top-K reference test to filter heap candidates now that
  exact maintained-value candidates also appear

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Sep 29, 2026
…472)

Logical-planning half of #462; no physical crate.

- Rate heap Top-K and grouped Rate/Sum summary placements (fixed-window
  and query-time Rate candidates, current-series/counter-value inputs)
- Exact-candidate admission: CostModel::summary_support_evidence hook,
  maintained exact values admitted alongside approximate Top-K/Count,
  division accepts exact guarantees
- Candidate-root enumeration (enumerate_candidate_dags[_for_root],
  assemble_selected_query) with budgeted, non-partial inventory
- PromQL series identity: reserved $promql_series_identity column so
  closed schemas can still be recognized as PromQL populations; direct
  rate/increase topk keeps counter-value ranking intent
- Per-series reduction resolves the sample column instead of guessing
- HLL confidence: standalone hll_confidence module dropped; it duplicated
  ClassicHllConfidence already on main in accuracy/estimators/hll.rs (#460)
- Update planner Top-K reference test to filter heap candidates now that
  exact maintained-value candidates also appear

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Sep 30, 2026
The #462 split no longer exposes PhysicalCandidate::encode; its serde form
gives the same byte-for-byte comparison.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Sep 30, 2026
The #462 split no longer exposes PhysicalCandidate::encode; its serde form
gives the same byte-for-byte comparison.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants